Papers with Elo ratings
Style Over Substance: Evaluation Biases for Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Ranking the relative performance of large language models based on Elo ratings is gaining popularity . however, the extent to which humans and LLMs are capable evaluators remains uncertain . |
| Approach: | They propose to evaluate machine-generated text across multiple dimensions using the Elo rating system . they propose to use crowd-sourced and expert annotators to rank models based on Elo ratings . |
| Outcome: | The proposed method improves the quality of LLM-based evaluations, but there is no improvement in crowd-sourced evaluations. |
War of Thoughts: Competition Stimulates Stronger Reasoning in Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have reshaped the landscape of reasoning tasks. |
| Approach: | They propose a method that enhances LLM reasoning without finetuning by using test-time scaling. |
| Outcome: | The proposed method outperforms baseline models in both budget and model size. |
PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments (2026.findings-acl)
Copied to clipboard
| Challenge: | evaluating therapeutic competence of large language models remains challenging due to unstructured and longitudinal nature of counseling. |
| Approach: | They propose a framework that calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments. |
| Outcome: | The proposed framework calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments. |